Papers with corpus creation methodology

2 papers
Twitter corpus of Resource-Scarce Languages for Sentiment Analysis and Multilingual Emoji Prediction (C18-1)

Copied to clipboard

Challenge: a majority of research studies on twitter focus on English tweets, despite the fact that English dominates the mix of languages.
Approach: They leverage social media platforms such as twitter for developing corpus across multiple languages . they use tweets to collect data for sentiment analysis and emoji prediction .
Outcome: The proposed method is applicable for resource-scarce languages provided speakers of that particular language are active users on social media platforms.
MuST-C: a Multilingual Speech Translation Corpus (N19-1)

Copied to clipboard

Challenge: Current research on spoken language translation (SLT) has to confront the scarcity of sizeable and publicly available training corpora.
Approach: They propose a multilingual speech translation corpus that will facilitate the training of end-to-end systems for SLT from English into 8 languages.
Outcome: The proposed multilingual speech translation corpus will facilitate the training of end-to-end systems for spoken language translation from English into 8 languages.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations